Back

DNA Research

Oxford University Press (OUP)

Preprints posted in the last 90 days, ranked by how well they match DNA Research's content profile, based on 26 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.

1
Chromosome-level genome assembly of macroalgae Gracilariopsis lemaneiformis

Hu, Y.; Huang, Y.; Yong, Y.; Shang, E.; Zhang, B.; Sui, Z.

2026-04-30 genomics 10.64898/2026.04.28.721235 medRxiv
Top 0.1%
18.4%
Show abstract

As an important cultivated red alga, Gracilariopsis lemaneiformis has great economic and ecological value. However, its existing genome assembly is highly fragmented and inadequately annotated. In this study, we constructed the first high-quality chromosome-level genome of Gp. lemaneiformis using PacBio long reads, Illumina short reads and Hi-C sequencing data. The assembled genome was approximately 86.66 Mb and the assembled sequences were anchored to 28 pseudo-chromosomes with lengths ranging from 1.70 to 7.81 Mb. 99.91% of the PacBio reads could be mapped to our assembly. In total, 8,664 genes were annotated, and the repeat elements identified in Gp. lemaneiformis constituted 65.04% of the whole genome, including 2.24% tandem repeat sequences and 62.81% interspersed repeats. We also established a high-evidence phylogenetic tree from 19 representative algae species, with the main aim to calculate their divergence times. This high-quality genome of Gp. lemaneiformis provides a crucial foundation for understanding genetic characteristics, investigating the genomic evolution, and facilitating molecular breeding.

2
A high-quality chromosome-scale reference genome assembly for Asparagus racemosus var. CIM-Shakti (Shatavari), a medicinal plant of Ayurvedic importance

Tyagi, S.; Sharma, A.; Shivani, K.; Gupta, V.; Paterson, A. H.; Trivedi, P. K.

2026-06-11 bioinformatics 10.64898/2026.06.07.730773 medRxiv
Top 0.1%
10.8%
Show abstract

Asparagus racemosus Wild., commonly known as Shatavari, is an important medicinal plant in Ayurveda and is valued for its steroidal saponins, particularly shatavarin compounds, which contribute to its adaptogenic, galactagogue, immunomodulatory, and therapeutic properties. Despite its medicinal and economic importance, genomic resources for this species have remained limited, restricting molecular breeding, pathway discovery, and comparative evolutionary studies within Asparagaceae. Here, we report a high quality chromosome scale reference genome assembly of A. racemosus var. CIM Shakti generated using PacBio HiFi long read sequencing and Omni C chromatin conformation scaffolding. The pseudo haploid assembly spans 817 Mb across 53 scaffolds, with a scaffold N50 of 98.50 Mb, L50 of 5, and a largest scaffold of 113.80 Mb. Ten major chromosome scale pseudomolecules were resolved, corresponding to the haploid chromosome complement of A. racemosus. The assembly showed high gene space completeness, with BUSCO completeness of 99.8% against the Eukaryota dataset and 98.0% against the Embryophyta dataset. BlobToolKit profiling further supported assembly quality, with GC content of approximately 39 to 40% and no major evidence of contamination. EDTA based repeat annotation identified 580.93 Mb of interspersed repetitive elements, accounting for 71.06% of the 817.57 Mb genome assembly. The repeat landscape was dominated by LTR retrotransposons, particularly Gypsy elements, which accounted for 25.01% of the assembly, followed by unclassified LTR elements at 26.58% and Copia elements at 4.84%. Structural and functional annotation identified 29,199 protein coding genes represented by 29,199 transcript models, 138,433 exons, and 125,201 CDS features. The annotation was structurally robust, with an average gene length of 4,605.1 bp, 4.74 exons per transcript, and 97.80% of transcripts containing multiple exons. The CIM Shakti reference genome provides a foundational genomic resource for investigating steroidal saponin biosynthesis, sex chromosome evolution, repeat driven genome expansion, and comparative genomics in Asparagaceae. This assembly will support future studies on medicinal trait improvement, conservation genomics, and genomics assisted breeding of climate resilient Shatavari cultivars.

3
Chromosome-level genome assembly of the Northeast China Brown Frog (Rana dybowskii)

zhang, y.; Wang, D.; Zhao, R.; Li, S.; Zheng, X.; Hu, G.

2026-06-15 genomics 10.64898/2026.06.11.731602 medRxiv
Top 0.1%
7.8%
Show abstract

Rana dybowskii is distributed across Northeast Asian and represents a valuable medical resource. A high-quality assembly of the genome has not yet been reproted. This species has 2n=24 chromosomes, but a huge genome size that estimated at 3.5 ~4.6 Gb in the previous studies. The relatively large chromosome size, exceeding hundreds of megabases, may result in difficulties of obtaining a complete chromosome level genome. Here, we constructed a chromosome-level genome assembly of R. dybowskii by integrating PacBio HiFi long-read sequencing for de novo assembly and CiFi (3C coupled with HiFi sequencing) for scaffolding. The final assembly consists of 12 chromosomes with a total of 3.95 Gb and a scaffold N50 length of 455 Mb. BUSCO assessment using the tetrapoda_odb12 database identified 94.2% complete and 0.5% fragmented orthologs, suggesting a high level of completeness of the assembly. Genomic annotation revealed that repetitive sequences comprise over 53% of the assembly, with retroelements and DNA transposons accounting for 22% and 25%, respectively. A total of 43,999 protein-coding genes were predicted with the assistance of RNA-seq reads from four tissues (muscle, eye, testis and skin). This high-quality chromosome-level reference genome provides a valuable genomic resource for advancing genetic studies of the species.

4
New chromosome-level haplotyped genome assemblies and annotation for the Japanese Quail (Coturnix Japonica)

Cabau, C.; Degalez, F.; Leroux, S.; Gourichon, D.; Serre, R.-F.; Vernette, C.; Donnadieu, C.; Iampietro, C.; Vandecasteele, C.; Pitel, F.; Klopp, C.

2026-05-14 genomics 10.64898/2026.05.12.724545 medRxiv
Top 0.1%
7.6%
Show abstract

The Japanese quail (Coturnix japonica) is a widely used model organism in developmental biology, genetics, and agriculture. Here, we present new, haplotyped, high-quality genome assemblies of the Japanese quail, generated using a combination of state-of-the-art sequencing technologies, including PacBio HiFi long reads, Oxford Nanopore sequencing, and Hi-C scaffolding. This assembly has a total length of 1.19 Gb, 80% of which is included in chromosomes, and is highly complete (BUSCO score aves_odb10: 97.3). Assembly metrics show a marked improvement in contiguity, with a significantly higher scaffold N50 and a lower number of contigs compared to the reference genome assembly. Remarkably, the assembly extends previously truncated chromosome ends, with 31 telomeres detected. In addition, we merged the existing Ensembl and Refseq annotations and obtained a combined set of 26,102 genes, of which 25,038 genes were successfully mapped on the improved assembly haplotype 1 (Cjap1.hap1). Together, these new genome assemblies and their enriched annotation provide a robust genomic framework for future research. They enhance our ability to investigate developmental processes, genetic and epigenetic inheritance, and host-pathogen interactions. Furthermore, they offer valuable insights for conservation genetics and sustainable breeding programs. This resource represents a critical step forward in leveraging the full potential of the Japanese quail as a model species in both basic and applied research.

5
Rubus armeniacus genome sequence reveals the secrets of blackberry anthocyanin biosynthesis

Wolff, K.; Nowak, M. S.; Thoben, C.; Beuerle, T.; Pucker, B.

2026-05-10 genomics 10.64898/2026.05.05.723051 medRxiv
Top 0.1%
5.6%
Show abstract

Here, we present a comprehensive multiomics analysis of anthocyanin biosynthesis in Rubus armeniacus, known for its dark fruits. A phased genome sequence of the tetraploid blackberry was generated, achieving an N50 of 34 Mb with an assembly size of 1.2 Gbp based on Oxford Nanopore Technology sequencing (ONT). The BUSCO score for the total assembly shows a high completeness of 99.1%. The assembly was separated into 4 pseudohaplophases, with the pseudohaplophase A representing the R. armeniacus genome in 7 chromosome scale contigs, with an N50 of 46 Mbp and 98.8% conserved BUSCO genes. A total of 118,183 protein coding genes were annotated within the genome assembly and all relevant genes encoding enzymes and transcriptional regulators of the anthocyanin biosynthesis pathway were identified within each pseudohaplophase. To further understand the underlying cause of dark pigmentation, the gene expression was analysed during different stages of berry development revealing a strong induction of anthocyanin biosynthesis genes including the anthocyanin activating subgroup 6 MYB transcriptions during the berry ripening process. Further, a quantification of cyanidin-3-O-glucoside in methanolic berry extract, utilizing a UHPLC-HRAM-MS analysis, revealed an approximately 500-fold increase of cyanidin-3-O-glucoside from green to black fruit, indicating that dark pigmentation in R. armeniacus results from high anthocyanin accumulation. Significance statementThis study provides a multiomics analysis of the dark pigmentation of Rubus armeniacus, including a high quality phased assembly and an in-depth analysis of the anthocyanin biosynthesis pathway. A transcriptional and metabolomic analysis revealed that dark berry pigmentation is caused by a high accumulation of cyanidin-3-O-glucoside during fruit ripening.

6
Chromosome-level genome assemblies of the red algae Porphyra dioica and Porphyra linearis

Morcillo, J.; D hondt, S.; Lipinska, A.; Bouckenooghe, S.; Noyen, L.; Van de Vloet, A.; Vranken, S.; Knoop, J.; Leliaert, F.; De Clerck, O.

2026-05-16 genomics 10.64898/2026.05.14.725108 medRxiv
Top 0.1%
5.5%
Show abstract

As one of the earliest-diverging multicellular eukaryotic lineages, the bladed Bangiales (Rhodophyta) possess a deep evolutionary history with a central role in the multi-billion-dollar global seaweed aquaculture industry. Although North Atlantic representatives are emerging candidates for regional mariculture, the scarcity of high-quality genomic resources for these taxa hinders both fundamental research and commercial optimization. To address this, we present the first chromosome-level genome assemblies for two native European species: Porphyra dioica (150.44 Mbp) and Porphyra linearis (95.22 Mbp). By integrating Oxford Nanopore Technologies (ONT) long-read sequencing with Hi-C proximity ligation, we generated highly contiguous nuclear genomes resolved into five chromosomes. Structural gene models were predicted through the BRAKER3 pipeline, identifying 12,548 and 10,382 protein-coding genes for P. dioica and P. linearis, respectively. Subsequent homology-based functional annotation characterized 57.4% and 59.8% of these predicted proteins. Supplemented by circularized organellar genomes, these reference genomes provide a critical framework for future research, enabling comparative studies of Atlantic-Pacific divergence and facilitating the development of selective breeding programs for sustainable European aquaculture.

7
Genome Assembly of the Green Revolution Wheat Cultivar Pavon 76 Establishes a Reference for CIMMYT-Derived Wheat

Eurmsirilerd, E.; Resendiz, M.; Lana, G.; Li, Q.; Nawmi, F.; Chang, R.; Zhang, W.; Wu, T. X.; Sun, Z.; Wang, L.; Lonardi, S.; Melonek, J.; Chater, J. M.; Jia, Z.

2026-06-05 genomics 10.64898/2026.06.04.730269 medRxiv
Top 0.1%
4.9%
Show abstract

Common wheat (Triticum aestivum L.) is a cornerstone of global food security. However, the standard reference genome, IWGSC RefSeq v2.1, is derived from wheat cultivar Chinese Spring, a historically important lineage that is not representative of modern cultivated wheat. In contrast, Pavon 76 is a CIMMYT Green Revolution wheat cultivar grown worldwide and is present in the pedigrees of numerous elite wheats, as well as in the genetic background of over 600 cytogenetic stocks used in studies ranging from recombination to male sterility. As no Pavon 76 reference has been reported, interpretation of these Pavon 76-derived stocks has relied on the genetically distant IWGSC RefSeq v2.1. To address this gap, we generated PacBio HiFi reads and Hi-C read pairs to assemble Pavon 76s genome. The final assembly comprises the expected 21 chromosomes, spans 15 Gb, and has an N50 of 710 Mb. Annotation identified 167,201 genes, of which 89.5% were functionally annotated. Comparative sequence and gene synteny analysis between Pavon 76 and IWGSC RefSeq v2.1 revealed significant structural divergence, establishing that IWGSC RefSeq v2.1 is not fully representative of modern wheats. This robust Pavon 76 assembly establishes the first chromosome-scale de novo assembly of a CIMMYT-derived wheat and provides a foundation for future genetic and genomic studies.

8
The Chromosome-Scale Genome of Phyllanthus niruri Reveals Candidate Genes and a Putative Biosynthetic Framework for Phyllanthin Formation.

Khushi, K.; Ganesh, A.; Sharma, A.; Ravindran, F.; Srinivasan, S.; Choudhary, B.

2026-05-16 genomics 10.64898/2026.05.15.725390 medRxiv
Top 0.1%
4.8%
Show abstract

Phyllanthus niruri (Phyllanthaceae) is a medicinally important herb known for producing phyllanthin, a bioactive dibenzylbutane lignan with reported hepatoprotective and antioxidant properties. However, the biosynthetic basis of phyllanthin production remains unresolved, largely due to the absence of a reference genome for the species. We report a Chromosome-Scale Assembly of P. niruri generated by integrating PacBio HiFi long reads and Illumina short reads, followed by reference-guided scaffolding against Phyllanthus cochinchinensis. The assembly has an L50 of 7 and 97.6% BUSCO completeness. Annotation predicted 19,254 protein-coding genes (91.1% functionally annotated), with phenylpropanoid biosynthesis emerging as the most enriched specialized-metabolism pathway in the genome. Using pathway-guided genome mining, structural similarity analysis, and comparative metabolic reconstruction, we propose a putative biosynthetic pathway for phyllanthin originating from the phenylpropanoid-lignan branch through secoisolariciresinol-like intermediates, followed by terminal O-methylation reactions. A total of 305 unique candidate genes associated with the proposed pathway were identified, including expanded families of dirigent proteins, peroxidases, secoisolariciresinol dehydrogenases, and O-methyltransferases. Comparative transcriptomic analyses across related Phyllanthus species further supported the proposed pathway through coordinated expression of lignan-associated genes and tissue-specific enrichment of O-methyltransferases. This work provides the first reference genome for P. niruri and a prioritized candidate gene set for functional characterization of phyllanthin biosynthesis.

9
A high-quality, chromosome-scale genome assembly of the shade-tolerant wild rice, Oryza granulata

Zhang, F.; Yang, Y.-h.; Li, W.; Shi, C.; Zhu, X.-g.; Gao, L.-z.

2026-05-01 bioinformatics 10.64898/2026.04.28.721348 medRxiv
Top 0.1%
4.2%
Show abstract

Oryza granulata Nees et Arn. ex Watt, a diploid wild rice (GG genome), possesses exceptional shade tolerance and is a key genetic resource for rice improvement. However, previous genome assemblies lacked continuity and completeness. Here we present a chromosome-scale reference genome of O. granulata using PacBio SMRT (113x), Hi-C (95x), and Illumina sequencing. The final assembly is ~764.24 Mb, with a scaffold N50 of ~59.32 Mb, and ~96.47% of the sequence anchored to 12 chromosomes. BUSCO completeness is ~98.6%. We annotated ~42,064 protein-coding genes, of which ~95.39% were functionally annotated, along with ~73.46% repetitive elements. The genome assembly and raw sequencing data are available at NGDC (PRJCA061980), NGDC GSA (CRA068332), and NGDC GWH (GWHISVE00000000.1). This high-quality genome will serve as a fundamental resource for evolutionary genomics, conservation biology, and breeding of shade-tolerant rice cultivars.

10
Chromosome-level genome assembly of Calotes wangi with dynamic colour variation

Qiu, X.; Wang, Y.; Wen, J.; Chen, Y.; Zhao, L.; Jian, J.; Yang, W.

2026-07-10 evolutionary biology 10.64898/2026.07.07.736949 medRxiv
Top 0.1%
4.2%
Show abstract

The Wangs garden lizard, Calotes wangi, is a widely distributed agamid species in Southern China and Northern Vietnam and exhibits pronounced colour variation and rapid body colour change. Despite increasing interest in the genomic basis of colour variation, chromosome-level genomic resources remain limited in agamid lizards. Here, we generated a chromosome-level reference genome of C. wangi using PacBio HiFi sequencing and Hi-C scaffolding. The final genome assembly was approximately 1.66 Gb in size and comprised 6 macrochromosomes and 11 microchromosomes, with a contig N50 of 110.09 Mb and 98.9% complete BUSCO genes. A total of 20,442 protein-coding genes were annotated. Comparative genomic analyses identified 297 significantly expanded gene families, with enriched functions associated with steroid metabolism, chromatin regulation, and epigenetic processes. This high-quality genome assembly provides an important genomic resource for future studies of colour variation, phenotypic plasticity, and evolutionary diversification in agamid lizards.

11
A telomere-to-telomere (T2T) pig genome assembly reveals Y chromosome diversity and structural variations of Wuzhishan pigs

Ren, Y.; Wang, F.; Li, X.; Liu, G.; Sun, R.; Zheng, X.; Zhang, Y.; Lin, R.; Lu, X.; Chen, L.; Xin, W.; Fei, Y.; Chao, Z.

2026-04-27 genomics 10.64898/2026.04.23.720499 medRxiv
Top 0.1%
4.2%
Show abstract

BackgroudWuzhishan (WZS) pigs are native to Hainan Province of China, and serve as both important agricultural resources and biomedical models. Although the published WZS pig genome (T2T-pig1.0) even achieving telomere-to telomere (T2T) completeness, substantial genetic diversity still exists within the same pig breed, another WZS pig genome named WZS-T2T was assembled in this study. ResultsMultiple sequencing data were used to assemble genome, and finally yielded a [~]2.68 Gb telomere-to-telomere genome, with N50 length [~]142.87 Mb, and annotated protein coding genes of 23,100. Compared to T2T-pig1.0, QV and BUSCO value was higher, and the Y chromosome (ChrY) length was longer in WZS-T2T than that of T2T-pig1.0. ChrY of two WZS pigs shared 11 genes, including sex differentiation-related genes of SHOX, PRKX, and DDX3X, and SRY; however, energy metabolism gene SLC25A4 and the macrophage-related receptor gene CSF2RA of ChrY were specific to WZS-T2T. An inversion SV on chromosome 10 with length [~]33.86 Mb was identified between two WZS pigs, and three proofs were proposed for proving the accuracy sequence orientation of WZS-T2T.The genetic diversity was consistent with LD decay speed in population different analysis. WZS pigs exhibited higher genetic diversity than other four pig populations (Tunchang pigs, Yuxi black pigs, Large White pig, and Duroc pigs) examined in this study, and presented slower LD decay compared to other four breeds. ConclusionsTherefore, WZS-T2T provided a higher-quality assembly, and potential advantages of both agricultural production and biomedical targets for WZS pigs.

12
Whole-Genome sequencing of Indigenous Withania somnifera accession and comparative cytochrome P450 phylogenomics

gupta, S.; Misra, P.; Singh, R.; Dhar, M. K.

2026-05-01 genomics 10.64898/2026.04.28.721529 medRxiv
Top 0.1%
4.0%
Show abstract

Cytochrome P450 monooxygenases (CYP450s) are key oxidative enzymes that diversify plant specialized metabolites and play a central role in the biosynthesis of bioactive withanolides in Withania somnifera (L.) Dunal. Despite their importance, genome-wide information on CYP450s in W. somnifera has remained elusive. Herein, the first high-quality genome assembly (2.2 Gb, scaffold N50: 47.4 kb) of an Indian W. somnifera cultivar was generated using a hybrid Oxford Nanopore-Illumina sequencing strategy. Comparative analysis with the NCBI reference genome revealed moderate SNP and indel variations, reflecting intraspecific genetic diversity. A comprehensive CYP450 catalog was established and analyzed phylogenomically across nine plant genomes, encompassing both withanolide-producing and non-producing Solanaceae and non-Solanaceae species. Unique CYP families (CYP450A, CYP1194, and CYP705A) were detected exclusively in W. somnifera, suggesting lineage-specific metabolic innovations, while Solanaceae-restricted (CYP82E/M) and absent (CYP81B, CYP6) lineages highlight taxonomic divergence. Across all analyzed genomes, 36 conserved CYP450 subfamilies, including triterpenoid-associated members, were identified, suggesting a shared oxidative framework adaptable to specialized metabolism. Moreover, potential candidate genes in the triterpenoid pathway, including CYP72A692_1, CYP72A560_4, CYP716A48, CYP724B2, and CYP51G1, were identified through phylogenetic integration with functionally validated triterpenoid-modifying enzymes from other plant species. Gene family evolution analysis further revealed contraction of monoterpenoid-related subfamilies (CYP76A), implying a metabolic shift toward triterpenoid specialization. The comprehensive genome assembly and CYPome of W. somnifera offer a valuable resource for functional characterization, evolutionary analysis, and the identification of genes underlying its specialized metabolism. Furthermore, the study advances our understanding of CYP450 diversity and evolution, revealing lineage-specific innovations, conserved subfamilies, and key candidate genes involved in triterpenoid biosynthesis. Together, these findings lay a foundation for future functional studies and pathway engineering aimed at optimizing the metabolic potential of this important medicinal plant.

13
Haplotype-resolved genome of autotetraploid alfalfa (Medicago sativa) Regen-SY27x uncovers large scale structural variation and resistance gene dynamics

Kaur, H.; Cameron, C. T.; Gomez, A.; Mudge, J.; Farmer, A.; Shannon, L. M.; Samac, D. A.

2026-05-05 genomics 10.64898/2026.05.01.722254 medRxiv
Top 0.1%
3.9%
Show abstract

Polyploid genome assembly presents unique challenges due to extensive heterozygosity and complex haplotype structure. We report a haplotype-resolved, chromosome-scale assembly of Regen-SY27x, a genotype of autotetraploid alfalfa (Medicago sativa), which is widely used for genetic modification because of its excellent regenerative capacity in tissue culture. Using PacBio HiFi long reads, Omni-C scaffolding, and linkage map guided phasing, we generated a 3.2 GB assembly comprising four haplotypes with high contiguity and completeness. Kmer-based validation confirmed accurate haplotype separation, while linkage map integration and dotplot analysis identified and corrected chimeric scaffolds. Gene annotation yielded 221,688 protein-coding genes, with more than 99% assigned to pseudochromosomes. Repetitive elements accounted for 62.7% of the genome, dominated by long terminal repeat retrotransposons and a high fraction of Helitrons. The spatial enrichment of Helitrons within gene-dense distal chromosome arms underscores their pivotal role as key drivers of genomic innovation and gene family expansion. We identified 3,696 nucleotide-binding leucine-rich repeat R genes, with Toll/interleukin-1 receptor-like and Rx-type subclasses forming large tandem clusters across haplotypes. Comparative analyses revealed strong macrosyntenic conservation among Regen-SY27x and the publicly available Chinese alfalfa genomes but extensive structural variation both within Regen-SY27x haplotypes and between Regen-SY27x and the Chinese genotypes with tens of thousands of duplications, inversions, and translocations detected. These results demonstrate that a single autotetraploid individual captures extensive structural diversity, but individuals from different populations vary greatly. The Regen-SY27x assembly provides a foundational genomic resource for investigating polyploid genome evolution and identifying genetic variation relevant to biological and agronomic improvement in alfalfa. Article SummaryThis study presents the first chromosome-scale, haplotype-resolved genome assembly of the US alfalfa germplasm, Regen-SY27x, a key alfalfa genotype used widely for genetic engineering. We integrated HiFi long reads, Omni-CTM scaffolding, and linkage map-guided phasing to reconstruct all four haplotypes of this complex autotetraploid. Our results identified 221,688 protein-coding genes and reveal immense intra-individual structural variations dominated by small duplications. This high-quality reference serves as a foundational tool for the alfalfa community, enabling researchers to link complex structural diversity with agronomic traits and further enhance the biotechnological potential of this essential forage crop.

14
Draft genome assembly of the Woolly bottlebrush (Greyia radlkoferi) using Pacbio long-read sequencing technology.

Molotsi, A. H.; Masebe, T.; Nesengani, L. T.; Mdyogolo, S.; Tshilate, T. S.; Smith, R. M.; Hlongwane, N.; Hadebe, S.; Mafokwane, T. M.; Mapholi, N.

2026-05-28 genomics 10.64898/2026.05.26.727846 medRxiv
Top 0.1%
3.3%
Show abstract

The Woolly bottlebrush (Greyia radlkoferi) is an indigenous South African plant known for its ornamental appeal and potential medicinal uses. It naturally grows on rocky hillsides and grasslands and highly resilient to drought, temperature fluctuations, and nutrient-poor soils. Its flavonoid-rich compounds with anti-tyrosinase activity support its traditional use to treat skin pigmentation disorders in humans. Despite its outstanding ecological and biochemical characteristics, no reference genome is available for Greyia radlkoferi. Therefore, this study aimed to generate the first draft genome of the Greyia Radlkofleri using PacBio Sequel IIe HiFi long read sequencing. A total of 56.07 Gb HiFi data was generated, providing a total genome coverage of 270X. The assembled genome size was 206Mb, with the longest scaffold being 13.9 Mb. The assembly statistics yielded a scaffold and contig N50 of 10.1 Mb, and an L50 of 9, and with an overall GC content of 34.9 %. The genome scope profile set at kmer = 17 indicated that the genome is triploid. The genome annotation predicted 17,804 protein-coding genes and 17,804 transcripts with an average gene length of 3,116.03 bp. This is the first draft genome of its kind for the Greyia genus and provides a foundation for future studies aimed at elucidating the genetic basis of its environmental resilience and the biosynthetic pathways underlying its medicinal properties.

15
A high-quality genome resource for Cercospora cf. flagellaris, a causal agent of Cercospora leaf blight of soybeans

Carver, Z. A.; Price, T.; Richards, J. K.; Doyle, V. P.

2026-06-17 genomics 10.64898/2026.06.16.732711 medRxiv
Top 0.1%
2.6%
Show abstract

A highly contiguous and complete reference genome of Cercospora cf. flagellaris, the causal agent of foliar disease on many plant hosts including Cercospora leaf blight of soybean, was assembled using a combination of PacBio and Illumina sequencing reads. The genome assembly is 33.72 Mb in length and consists of 14 nuclear scaffolds and one mitochondrial contig. Four scaffolds have telomeric repeats on both ends and represent fully assembled chromosomes, while nine scaffolds represent partially assembled chromosomes with telomeric repeats on one end. The assembly has an N50 of 2.90 Mb and an L50 of 5 scaffolds. Genome annotation identified 11,268 genes, of which 947 and 360 were predicted to encode secreted proteins and effectors, respectively. Additionally, 512 genes were predicted to encode carbohydrate-active enzymes and 60 biosynthetic gene clusters were annotated. Taken together, this annotated genome assembly will be a valuable resource for genomics, host-pathogen interactions, and population biology research in this economically important pathosystem.

16
A Highly Contiguous Reference Genome for Scalesia gordilloi (Asteraceae), a Critically Endangered Plant Endemic to the Galapagos Islands

Pozo, G.; Rivas-Torres, G.; Velez-Darquea, E.; Barragan-Orbe, D.; Torres, M. d. L.

2026-06-29 genomics 10.64898/2026.06.25.734018 medRxiv
Top 0.2%
1.9%
Show abstract

Scalesia gordilloi is a critically endangered species endemic to San Cristobal Island in the Galapagos archipelago and represents one of the most unique and vulnerable lineages within the adaptive radiation of the genus Scalesia. Despite its evolutionary distinctiveness and conservation importance, no genomic resources have been available for this species. Here, we present the first high-quality reference genome of S. gordilloi, generated using Oxford Nanopore long-read sequencing. Across three PromethION R10.4.1 flow cells, we obtained 80.5 Gb of long reads (~25X coverage), which enabled a highly contiguous 3.61 Gb assembly composed of only 549 contigs and an N50 of 106.6 Mb. BUSCO completeness reached 98.6%, with assembly metrics comparable to other high-quality Asteraceae genomes. Repeat annotation revealed that 76.2% of the genome is composed of interspersed elements, dominated by LTR retrotransposons. Structural annotation resulted in 47,913 high-confidence protein-coding genes, consistent with expectations for large, repetitive Asteraceae genomes. This genome provides a critical foundation for conservation genomics, enabling assessments of genetic diversity, inbreeding, and adaptive potential in the species. It further establishes a framework for comparative genomics across the Scalesia radiation and supports future efforts to protect and restore one of the most threatened plant lineages of the Galapagos Islands.

17
A Draft Male Genome Assembly of the Slipper Lobster (Thenus australiensis) Reveals an XY System and a Validated Diagnostic Marker for Monosex Aquaculture.

Tran Nguyen, A. H.; Ha, G.-H.; Tran, D.-P.; Le, N. T.; Glendining, S.; Fitzgibbon, Q.; Herzig, V.; Luu, P.-L.; Ventura, T.

2026-06-29 genomics 10.64898/2026.06.24.734161 medRxiv
Top 0.2%
1.9%
Show abstract

The slipper lobster (Thenus australiensis) is rapidly emerging as a high-potential species for commercial aquaculture. Because females exhibit superior growth characteristics due to less frequent moulting after sexual maturity, developing monosex breeding strategies is highly desirable for industry profitability. However, the lack of genomic resources and early sex-identification tools has hindered this development. Here, we report the first draft male genome assembly for T. australiensis, generated using a combination of whole-genome shotgun sequencing, DArT-seq, and multi-tissue transcriptomics. The curated assembly spans 0.913 Gbp with high functional completeness (93.0% BUSCO), providing a robust repertoire of 30,100 protein-coding genes. Through k-mer subtraction and population-level DArT-seq genotyping, we provide definitive evidence that T. australiensis utilizes an XX/XY sex-determination system. Crucially, by identifying male-specific structural variations within a neo-Y locus, we developed a diagnostic PCR assay targeting a male-exclusive sequence. This 171 bp marker achieved 100% accuracy in phenotypic sex identification across wild-caught populations. Ultimately, these foundational genomic resources, combined with a highly reliable molecular sexing tool, provide the critical framework necessary for early sex sorting, broodstock management, and the commercial advancement of monosex slipper lobster farming.

18
LsBADH1 is responsible for sweet fragrance in lettuce (Lactuca sativa L.) through 2-acetyl-1-pyrroline biosynthesis

SEKI, K.; Matsui, K.; YANAGIDATE, M.; NISHIDA, K.; KOYAMA, R.; Uno, Y.

2026-06-16 genetics 10.64898/2026.06.13.731611 medRxiv
Top 0.2%
1.7%
Show abstract

HighlightThe sweet fragrance of lettuce was attributed, for the first time, to the synthesis of 2-acetyl-1-pyrroline caused by a deficiency in the betaine aldehyde dehydrogenase gene. Fragrance is among the most valuable traits of high-quality crops and influences consumer preferences. Although 2-acetyl-1-pyrroline (2AP) is a key component of fragrant cultivars in several crops, its genetic mechanism in lettuce (Lactuca sativa L.) remains poorly understood. The betaine aldehyde dehydrogenase (BADH) gene has been identified as causative for 2AP-derived fragrance in rice and soybean cultivars. Hence, we conducted a linkage analysis using an F2 population derived from a cross between Kukichisya (fragrant) and Rennet (non-fragrant) for three candidate genes of BADH orthologs in the lettuce genome. Analysis linked LOC111877932 located in LG4 to the fragrance trait, and it was designated LsBADH1. Comparison among Kukichisya, Salinas, and candidate BADH of sunflower (Helianthus annuus L.) revealed three non-synonymous single-nucleotide polymorphisms (nsSNPs) in exons 1, 2, and 9, and suggested that nsSNP in exon 9 was strongly correlated with fragrance in Kukichisya. A premature stop codon introduced in exon 5 of LsBADH1 using Target-AID base-editing technology resulted in truncated BADH1 and higher 2AP levels. Our results indicated that LsBADH1 is responsible for the 2AP-derived fragrance. Our findings can be applied to select cultivars based on a novel concept for the cooking process, providing a transformative platform to breed fragrant lettuce as a high-value-added product.

19
First high-quality genome assemblies with chromosome-scale contiguity of Tunisian durum wheat (Triticum turgidum subsp. durum) landraces Chili and Mahmoudi

GDOURA BEN AMOR, M.; MATHLOUTHI, N. E. H.; BELGUITH, I.

2026-05-28 genomics 10.64898/2026.05.25.727644 medRxiv
Top 0.2%
1.6%
Show abstract

Durum wheat (Triticum turgidum subsp. durum) is a globally important crop for pasta and couscous production. Chili and Mahmoudi are historically significant Tunisian landraces valued for exceptional grain quality, high protein content, and adaptation to arid Mediterranean climates. Yet no high-quality reference genome assemblies were available for either variety before this work. We assembled both genomes using publicly available PacBio HiFi long reads and Illumina Hi-C proximity ligation data deposited under NCBI BioProject PRJNA1420514. HiFi reads were assembled with hifiasm v0.25.0 in primary mode, and Hi-C scaffolding was performed with YAHS v1.2a.2 after read alignment with BWA-MEM. Assembly quality was assessed with QUAST v5.3.0 and BUSCO v5.8.0 (embryophyta_odb10 lineage). The Chili assembly spans 10.84 Gbp across 3,472 scaffolds with a scaffold N50 of 844.9 Mbp and BUSCO completeness of 99.4%. The Mahmoudi assembly spans 10.70 Gbp across 3,258 scaffolds with a scaffold N50 of 2,072 Mbp and BUSCO completeness of 99.3%. Merqury v1.3 confirmed high base accuracy (QV 68.0 for Chili, QV 68.3 for Mahmoudi) and k-mer completeness (>98% for both). Independent validation with wfmash confirmed 98.6% mean alignment identity across 11,172 chromosome-to-reference alignments. Post-assembly characterization of GC profiling, centromere architecture, ribosomal DNA arrays, and structural variation revealed extensive genome-level detail. Both assemblies substantially exceed the contiguity of existing durum wheat references and represent the first chromosome-scale-contiguity genomic resources for North African durum wheat landraces. The Mahmoudi accession carried a putative 2B-3B homeologous fusion on chromosome 3B (4,710 Mbp), a structural novelty in an ancient Tunisian landrace. Assembly and Pseudomolecules available from Zenodo (10.5281/zenodo.20366290). The workflow was executed reproducibly on the public Galaxy Europe platform, demonstrating that reference-quality plant genome assembly is achievable without local HPC infrastructure. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=179 HEIGHT=200 SRC="FIGDIR/small/727644v2_ufig1.gif" ALT="Figure 1"> View larger version (52K): org.highwire.dtl.DTLVardef@159c9d4org.highwire.dtl.DTLVardef@1d1a0deorg.highwire.dtl.DTLVardef@19853dborg.highwire.dtl.DTLVardef@1a96018_HPS_FORMAT_FIGEXP M_FIG C_FIG HighlightsO_LIFirst chromosome-scale-contiguity assemblies for Chili and Mahmoudi landraces C_LIO_LIChili genome: 10.84 Gbp, scaffold N50 844.9 Mbp, BUSCO 99.4% C_LIO_LIMahmoudi genome: 10.70 Gbp, scaffold N50 2,072 Mbp, BUSCO 99.3% C_LIO_LIMerqury QV ~68 confirms base accuracy, >98% k-mer completeness C_LIO_LI>140-fold scaffold N50 improvement over Svevo v1 reference C_LI

20
Evolutionary dynamics of Aegilops revealed through comparative genome assembly of all 25 species

Shazadee, H.; Edwards, T.; Levesque-Lemay, M.; Zheng, C.; Ens, J.; Pozniak, C. J.; You, F. M.; Cloutier, S.

2026-07-10 genomics 10.64898/2026.07.09.737531 medRxiv
Top 0.2%
1.5%
Show abstract

Aegilops species are the closest wild relatives of wheat and an important reservoir of genetic diversity for its improvement. Despite their potential, many Aegilops genomes remain poorly characterized. Here we present high-quality assemblies of 18 diploid, tetraploid, and hexaploid Aegilops genomes, which, along with the previously published genomes, complete the production of reference assemblies for all 25 genomes in this genus. Assembly sizes ranged from 5.24 Gb in diploids to 12.65 Gb in hexaploids, with scaffold N50 values up to 749.2 Mb. Gene annotation identified 53,035-156,779 protein-coding genes, of which 21,865-60,490 were classified as high-confidence. Orthogroup-based pangenome analysis across the 25 Aegilops genomes identified 80,521 orthogroups, including 15,809 core, 61,735 dispensable, and 2,977 species-specific orthogroups, highlighting substantial gene content variation among genomes. Phylogenetic analysis of 63 Triticum and Aegilops genomes/subgenomes based on near single-copy orthologs defines the phylogenetic relationships within the Triticum/Aegilops complex and confirms diploid progenitors of polyploid lineages. Ae. mutica (T) and Ae. speltoides (S) belong to the B lineage while the remaining Sitopsis grouped within the D lineage. Structural variation analyses using diploid progenitors as references revealed extensive large-scale rearrangements following polyploidization, emphasizing the dynamics of their evolution. Transposable element (TE) annotation further highlighted subgenome-specific TE expansions and contractions, providing insights into the mechanisms shaping genome structure after polyploidization. Collectively, these genomic resources provide a comprehensive framework for exploring Aegilops diversity, understanding polyploid evolution, and accelerating wheat improvement.